Back

BMJ Health & Care Informatics

BMJ

Preprints posted in the last 90 days, ranked by how well they match BMJ Health & Care Informatics's content profile, based on 15 papers previously published here. The average preprint has a 0.04% match score for this journal, so anything above that is already an above-average fit.

1
Are Frontier Large Language Models Safer Than Government-Backed Symptom Checkers for Clinical Self-Triage? A Standardised Vignette Evaluation

Chowdhury, A. R.; Chowdhury, B.

2026-09-02 health informatics 10.64898/2026.09.01.26361908 medRxiv
Top 0.1%
12.8%
Show abstract

Background: Consumer use of AI chatbots for health advice is rising, yet triage safety relative to established services remains unclear. Australia's Healthdirect, a government-backed symptom checker with 2.4 million uses in FY2024-25, remains unevaluated against frontier large language models (LLMs), and whether premium subscriptions improve triage safety remains unexplored. This study compared the triage accuracy and safety of Healthdirect against six LLM configurations across ChatGPT, Claude, and Gemini, assessed whether paid subscriptions improve triage safety, and characterised each system's error patterns. Methods: Forty-five clinical vignettes from the Semigran et al. benchmark spanning emergency, non-emergent, and self-care categories (15 each) were evaluated across seven systems. Healthdirect was tested following a seven-rule interaction protocol. LLMs were evaluated using first-person patient-language prompts under free-tier and paid-tier conditions. Outcomes were triage accuracy, emergency sensitivity, under-triage, and critical misses, analysed using Cochran's Q, Bonferroni-corrected McNemar tests, Cohen's kappa, and Wilson intervals. Findings: Triage accuracy differed significantly (Cochran's Q = 36.79, p < 0.001). Healthdirect achieved 48.9% accuracy (95% CI 35.0% to 63.0%; kappa = 0.233) versus 73.3% to 86.7% for LLMs (kappa = 0.600 to 0.800). Healthdirect operated under conservative interactive defaults while LLMs received complete information in a single prompt, which may have disadvantaged Healthdirect. Emergency sensitivity was 46.7% versus 80.0% to 86.7% for LLMs. Healthdirect produced two critical misses; no LLM produced any across 270 evaluations (95% CI 0% to 1.4%). When LLMs undertriaged, they recommended GP care rather than self-care. No tier differences were significant (all p > 0.05), and most systems over-triaged self-care cases. Interpretation: Frontier LLMs demonstrated higher triage accuracy and safer error profiles than Healthdirect. All LLMs avoided critical misses; Healthdirect did not. Premium subscriptions did not significantly improve triage safety. These findings support clinical governance decisions about whether LLMs warrant formal evaluation alongside government-backed symptom checkers.

2
Community Learning Ledgers for Cancer Navigation in Small Island Developing States

Roach, A.; Amow, A.; Haraksingh, R.; Archer, N.; Cyrus, E.; Evans, A. N.; Calleja, N.; Croes, R.; Forghani, I.; Bajnath, A.; Hadley, D.

2026-08-18 health informatics 10.64898/2026.08.16.26360547 medRxiv
Top 0.1%
11.7%
Show abstract

Importance. Cancer is the second leading cause of death among patients in the Caribbean, where outcomes are associated with delayed clinical navigation to screening, diagnosis, and treatment. Artificial intelligence is increasingly used to guide patients with cancer to care, but whether these systems provide clinically actionable, facility-verified guidance for individuals in this population, and whether governance of the system is associated with the quality of that guidance, has not been evaluated. Objective. We tested whether a governed community learning platform navigates Caribbean cancer patients better than four ungoverned AI systems, and we tracked how community intelligence accumulates over time. Design, Setting, and Participants. We deployed a community learning ledger (CaribChat.ai) across ten Caribbean jurisdictions beginning March 2, 2026, and report all sessions through June 1, 2026 (N=207). An initial actively-promoted accrual period (March 2 - April 6, 2026; 168 sessions) was followed by continued organic use after active clinical promotion ceased. We then submitted the same 28 patient screening queries to ChatGPT (GPT-4o), Claude Haiku 4.5, DeepSeek-Chat, and OpenEvidence on April 5-6, 2026. Claude Haiku 4.5 powers CaribChat; testing it without governance isolates the governance effect. The platform requires no registration. Exempt under 45 CFR 46.104(d)(4)(ii). Main Outcomes and Measures. We classified 207 community sessions by thematic domain and temporal phase. We scored each of five systems on Caribbean facility citation, actionable navigation, and US-resource leakage across 28 screening queries. Results. The ledger accumulated 207 sessions - 168 during an actively-promoted accrual period (March 2 - April 6) and 39 after active clinical promotion ceased. Community engagement evolved from screening questions to active treatment navigation and diaspora engagement. CaribChat cited verified Caribbean facilities in 28/28 (100%) responses versus 10/28 (35.7%) for ChatGPT and 9/28 (32.1%) for OpenEvidence. CaribChat provided actionable navigation in 28/28 (100%) versus 2/28 (7.1%) for OpenEvidence (P<=.001). The same model scored 100% with governance and 54% without (P<=.001). DeepSeek cited US resources in 57.1% of Caribbean responses. After active clinical promotion ceased, off-codebook queries rose from 2.4% to 26.7% across phases while the governance contract continued to reject every adversarial probe - the community persisted but drifted from the cancer codebook absent clinician curation. The deployment operated within the OECS Health Strategy 2030 and CARICOM regional health frameworks, with queries originating across Caribbean jurisdictions led by Trinidad and Tobago. Conclusions and Relevance. Every ungoverned AI system we tested failed Caribbean cancer navigation. The best scored 68%. The most widely adopted physician platform scored 7%. The same foundation model scored 100% with governance and 54% without. Community intelligence accumulated from the population it serves, not published literature, is what makes health AI work in SIDS. The post-promotion decay shows the requirement is bidirectional: sustained, on-codebook engagement depends on patients and clinicians working together - community participation and active clinical curation are jointly necessary for maximum AI leverage.

3
Computer Vision for Real-Time Anatomical Navigation in Neurosurgery: First-in-Human Clinical Evaluation and Iterative Development (IDEAL Stage 1)

Khan, D. Z.; Mao, Z.; Wijekoon, A.; Das, A.; Williams, S. C.; Blandford, A.; Jain, A.; Harris, L.; Borg, A.; Dorward, N. L.; Clarkson, M.; Bano, S.; McCulloch, P.; Stoyanov, D.; Marcus, H.

2026-06-11 surgery 10.64898/2026.06.11.26355205 medRxiv
Top 0.1%
11.0%
Show abstract

Introduction: Precise anatomical navigation is fundamental to safe endoscopic pituitary surgery, a high-stakes procedure characterised by a challenging learning curve. While traditional navigation systems often rely on workflow-disrupting probes or static preoperative imaging, advancements in computer vision AI (CVAI) now enable dynamic, real-time anatomical segmentation directly from live surgical video1-3. Our group has previously conducted a series of preclinical human-computer interaction studies to refine the system's design, alongside digital and high-fidelity physical simulations demonstrating the benefit of AI assistance in improving overall performance, training, and safety4-8. Building on this foundation, the current study represents a first-in-human application of real-time CVAI assistance in the neurosurgical operating room, serving to assess feasibility and safety, and to iteratively improve the system. Method: Guided by DECIDE-AI and IDEAL frameworks, this single-centre evaluation comprises an initial proof-of-concept phase (n=6) for endoscopic transsphenoidal pituitary surgeries. The AI model utilised a DINOv3-derived vision transformer architecture, deployed via a high-performance edge computing unit to achieve low-latency, real-time inference without reliance on cloud infrastructure2. Given the high-risk nature of the procedure and the early stage of clinical AI integration, the system was initially deployed as an educational adjunct on a secondary monitor, ensuring the primary surgical feed remains uncompromised. Functionality and safety were assessed via structured questionnaire, prospective observation, and blinded retrospective review of the recordings of the endoscopic surgical video feed and wider operating room environment. Continuous multi-stakeholder feedback through validated human factors surveys drove iterative technical refinements between cases. Results: Six patients with pituitary adenomas were enrolled. The CVAI system was successfully deployed in four cases, demonstrating acceptable real-time sella segmentation accuracy. Deployment failed pre-operatively in two cases owing to a single recurring system reboot bug. Iterative refinement between cases were driven by our experience and surgical team feedback. This resulted in the integration of additional anatomical structure segmentations (e.g., carotid arteries), enhanced model accuracy via training dataset expansion, and hardware firmware upgrades. Multi-stakeholder surveys demonstrated satisfactory system feasibility, usability, and acceptability among the surgical team. Both prospective observation and retrospective video review confirmed the absence of adverse events, including no significant distraction to the primary surgeon, and there were no AI-related clinical complications. Conclusion: This first-in-human early clinical evaluation demonstrates the feasibility, safety and iterative development of real-time, CVAI-based anatomical navigation during high-stakes neurosurgery. Future work will include a larger single-centre case series (IDEAL Stage 2a) with more surgical teams to further iterate the system and explore its impact on training and workflow. As the underpinning technology improves, deployment will transition to direct intra-operative decision support and integration with other intra-operative navigational technologies.

4
Development of an interdisciplinary network to improve the capacity to conduct digital legacy research: a quality improvement initiative

Nwosu, A. C.; Tibbles, A.; Goodwin, C.; Kaye, L.; Stanley, S.

2026-08-10 palliative medicine 10.64898/2026.08.06.26359876 medRxiv
Top 0.1%
10.3%
Show abstract

Background Digital legacy (the digital information available about someone following their death) has increasing societal importance as personal assets and interactions become increasingly digitized. Healthcare professionals often have a limited understanding of how to address digital legacy in practice, and there is a lack of interdisciplinary networks to improve education, research, and professional development in digital legacy. Objective This paper describes the development of an interdisciplinary initiative designed to build research capacity and develop consensus-based recommendations for integrating digital legacy into palliative care. Method Over 12-months, we conducted interdisciplinary engagement activities with diverse stakeholders, including clinicians, designers, and sociologists. We used a modified World Cafe method to facilitate dialogue and capture feedback on how memories are digitally curated, the management of digital estates, and intergenerational perspectives on digital legacy. Results We identified eight core recommendations for research and policy, including promoting digital legacy education, supporting policy development, and broadening the scope of interdisciplinary research. Our discussions highlighted the complexity of modern digital estates and the need for legal and ethical frameworks to protect individual rights. Conclusions The Network demonstrates that interdisciplinary collaboratives can address important issues relating to digital legacy, which provides a foundation to conduct collaborative research that improves the management of digital legacies in society.

5
Multisite Real-World Validation of an Electronic Health Record-Integrated Generative Artificial Intelligence Tool for Venous Thromboembolism Risk Stratification

Baughman, D. J.; Liu, S.; Jee, S.; Young, C.; Knight, A. M.; Davis, A.; Yegnasubramanian, S.; Najjar, P.; Whitbread, J. J.; Ahumada, L.; Chused, A.; Haut, E. R.; Lau, B. D.; Sridharan, A.; Streiff, M.; Aziz, K. B.

2026-06-22 health informatics 10.64898/2026.06.17.26355819 medRxiv
Top 0.1%
9.7%
Show abstract

Background: Guiding risk-appropriate inpatient thromboprophylaxis requires venous thromboembolism (VTE) risk stratification; however, reliable risk determination remains inconsistent in routine care. Health systems increasingly pilot artificial intelligence (AI) tools, yet few studies demonstrate rigorous evaluation in the context of a learning health system (LHS). We evaluated the performance of a pilot electronic health record (EHR)-integrated generative AI (GenAI) system, inHealth General Reasoner (iHGR), for VTE risk stratification versus clinician order set classifications and physician-adjudicated chart review. Methods: This multisite retrospective validation study included adult inpatient admissions at Johns Hopkins Medicine between June 21, 2025, and Dec 18, 2025 (checklist-based order set from June 21, 2025 - November 19, 2025, and clinician judgement-based order set from November 29 - December 18, 2025). From 758 eligible admissions, we randomly sampled 500 balanced by site and order set periods. iHGR and clinician-selected order set classifications were compared with the reference standard (RS). Primary outcomes were iHGR sensitivity and specificity. Secondary analyses compared the order sets with the same RS to evaluate workflow comparators and error patterns. Results: iHGR achieved 81.8% sensitivity (95% CI 77.3-85.6) and 70.9% specificity (63.6-77.3). The checklist-based order set had 61.3% sensitivity (53.7-68.5) and 86.2% specificity (77.4-91.9). The clinician judgement-based order set had 78.1% sensitivity (71.3-83.7) and 65.4% specificity (54.3-75.0). False-negative iHGR classifications were associated with missed narrative risk factors. Conclusion: iHGR showed higher sensitivity for VTE risk than checklist-based order sets and clinician judgement without introducing systematic bias. In silico evaluation of pilot AI systems within LHSs can identify clinically important performance trade-offs and implementation targets before operational scale-up. Narrative clinical data abstraction remained a key limitation, supporting the use of GenAI to support rather than supplant clinician judgement.

6
Development and formative application of the Health Data Readiness Level framework for federated health-data services.

Seymour, D.; Halliday, R.; Smart, J.; Burns, F.

2026-07-27 health informatics 10.64898/2026.07.23.26358713 medRxiv
Top 0.1%
9.7%
Show abstract

Objectives: To develop a multidimensional framework for assessing organisational and system readiness for federated health-data services, and to report its formative application across three heterogeneous UK ecosystems. Methods: The Health Data Readiness Level (HDRL) framework was developed from a structured landscape review of 56 maturity and readiness frameworks, first-principles requirements analysis, artificial-intelligence-assisted synthesis with human source verification, and stakeholder refinement. It comprises 64 indicators in eight domains and five ordered levels. Formative application examined three heterogeneous UK health-data research ecosystems using documentary evidence, professional stakeholder input, workshops, a structured right-of-reply process, cross-case calibration, and two illustrative research use cases. Analysis was descriptive; the application was not designed as psychometric validation or a league table. Results: All 64 indicators were scoreable in each case. Publicly reported profiles ranged from Developing (Level 2 -3) to Managed (Level 3 -4); all three cases met the proposed minimum for five foundational indicators. Recurring constraints concerned evidence of measured service performance, national-scale primary-care data access, cross-jurisdiction governance reciprocity, workforce capacity, and sustainable funding. Pandemic-era four-nation research was delivered through coordinated local analyses and meta-analysis, rather than routine automated federation. Discussion: HDRL operationalises a broad service- and system-readiness view that complements technical and governance specifications. Content validity, inter-rater reliability, aggregation choices, responsiveness and predictive validity remain to be established. Conclusion: HDRL is an evidence-informed candidate improvement and planning instrument for federated health-data services. It should not yet be used as an accreditation standard or as an official participation threshold for the UK's Health Data Research Service.

7
Quality and Safety profiles of AI-Generated vs Clinician-Generated Handoffs in Hospital Medicine

Shah, K. P.; Airan Javia, S.; Savage, T.; Bressman, E.

2026-06-08 health informatics 10.64898/2026.06.05.26354946 medRxiv
Top 0.1%
9.5%
Show abstract

End-of-rotation handoffs are critical for patient safety but add to documentation burden for hospitalists. Generative artificial intelligence (AI) may help automate handoff creation using electronic health record data, but its impact on quality and safety is unclear. Methods: We developed an AI handoff tool with a large language model using clinical notes as input and conducted a retrospective evaluation comparing AI-generated and clinician-authored handoffs. Handoffs were assessed across domains of quality and safety through a structured review. Results: Quality ratings were similar between AI and human handoffs (3.7 vs. 3.5, p=0.57). AI-generated handoffs were rated higher for organization (4.4 vs. 4.1, p=0.05) and completeness (4.1 vs. 3.6, p=0.01), but lower for conciseness (3.7 vs. 4.1, p=0.03) and accuracy (4.1 vs. 4.4, p=0.03). Error rates were comparable (0.3/handoff in both groups); however, AI-generated handoffs included inaccuracies (9% of AI errors) and hallucinations (1% of AI errors), while clinician-authored handoffs contained only omissions. Conclusion: Human and AI handoffs have differing error profiles and tradeoffs between completeness and conciseness. Prospective evaluation in clinical workflows is underway.

8
Neuro-Symbolic AI for Automated Pathology Quality Measurement

Brann, F.; Tadele, L.; Skau, C.; Bocsi, G.; Clarke, A. K.

2026-07-24 health systems and quality improvement 10.64898/2026.07.22.26358635 medRxiv
Top 0.1%
7.9%
Show abstract

Background. Clinical quality measurement often relies on manual abstraction of medical records, an approach that is costly, burdensome, and often infeasible for measures requiring interpretation of narrative text; these constraints have shaped measure development itself, filtering out clinically important measures that are too difficult to operationalize. We evaluated whether neuro-symbolic artificial intelligence (NSAI), which combines large language model extraction with symbolic reasoning, could reliably abstract complex quality measures from narrative pathology reports. Methods. The NSAI system decomposes each measure into atomic questions and is aligned to real-world reports through case-based refinement, an iterative human-in-the-loop process. Using 2,000 independently double-abstracted reports, we compared NSAI-based abstraction against trained human abstractors across four pathology quality measures established by the College of American Pathologists. Results. The NSAI system's agreement with the adjudicated gold standard (Cohen's kappa = 0.95) matched or modestly exceeded that of the trained human abstractors measured against the same standard (kappa = 0.92), with particularly strong performance on Gastrointestinal Metaplasia (CAP 43). In component analyses, case-based refinement drove the largest accuracy gains (up to 0.25 in kappa), whereas architectural decomposition primarily reduced performance variance across language-model backends more than tenfold, a property essential for clinical deployment. Conclusions. These findings suggest that automated abstraction could enable census-level quality measurement, reduce reporting burden, and expand the range of clinically meaningful measures that can be operationalized from narrative clinical documentation.

9
Entity-Aware Generation of Synthetic Clinical Progress Notes for Prostate Cancer using Large Language Model

Rey-Blanes, A.; Veredas-Morente, J.; Moreno-Barea, F. J.; Veredas, F. J.

2026-06-15 health informatics 10.64898/2026.06.12.26355166 medRxiv
Top 0.1%
7.8%
Show abstract

Objectives: This study investigates large language models (LLMs) for clinical entity projection across substantial textual transformation. Specifically, we evaluate whether entities annotated in Spanish prostate cancer case reports can be preserved and explicitly projected when the source narratives are transformed into hospital-style clinical progress notes. Entity projection is treated as a generation-driven task, allowing paraphrase, condensation and narrative reorganisation, providing that clinically relevant entities remain recoverable as structured annotations. Methods: A corpus of 109 Spanish prostate cancer case reports was annotated using a silver-standard pipeline combining Spanish biomedical named-entity recognition with rule-based prostate-specific antigen (PSA) and Gleason extractors. The resulting silver-standard annotations were validated on a subset of generated notes against a gold-standard consensus produced by medical experts in prostate cancer. Four LLMs were evaluated for note generation and entity projection: GPT-5.4 Nano, Qwen 3.5:35B-A3B, GLM5 and Claude Sonnet 4.6. Entity-to-Entity (E2E) generation used XML-annotated cases as RAG-supported input, whereas Text-to-Entity (T2E) generation required models to generate and annotate notes directly from plain text cases. Zero-shot and few-shot prompting were tested. Projection quality was measured using precision, recall and F1-score, and complemented by LLM-as-a-judge evaluation using Kimi K2.6. Results: E2E consistently outperformed T2E, indicating that explicit entity-enriched in- put substantially facilitates entity preservation and localisation. GLM5 achieved the best E2E zero-shot result (F1 = 0.915), followed by Claude Sonnet 4.6 (F1 = 0.896). In T2E, few-shot prompting improved performance, with Claude Sonnet 4.6 reaching the highest score (F1 =0.718). Age, Gleason, Disease, Procedure, Duration and negation-related entities were robustly projected, whereas PSA and Dose showed less stable behaviour. Conclusion: LLMs can generate clinically plausible synthetic prostate cancer evolution notes while preserving a substantial proportion of source entities, particularly when explicit semantic annotations are provided as input. However, the lower and more variable performance observed in T2E highlights the difficulty of jointly generating clinical narratives and projecting entities without source-side information, especially for numerical and measure-related entities.

10
Comparative analysis of discriminative and generative natural language processing pipelines for automated prostate magnetic resonance imaging reports

Lee, D. J.; McCoy, N.; Haroldsen, C.; Gilkey, M.; Verma, S.; Pyarajan, S.; Maxwell, K.; Nickols, N.; Rettig, M.; Silvestri, G.; Garraway, I.

2026-07-14 health informatics 10.64898/2026.07.12.26357886 medRxiv
Top 0.1%
7.6%
Show abstract

Objectives: Natural language processing (NLP) can enable scalable extraction of clinically relevant information from unstructured radiology reports retrieved from electronic healthcare data warehouses, but reliance on externally hosted models may pose cost, privacy, and deployment challenges. We compared self-hosted discriminative and generative NLP pipelines for automated extraction of Prostate Imaging and Reporting Data System (PIRADS) scores from multiparametric magnetic resonance imaging (mpMRI) reports used in prostate cancer risk assessment. Materials and Methods: We identified 44,511 mpMRI reports across 68 Veterans Affairs (VA) healthcare systems. A stratified random sample of 1,973 reports was used to train, test, and evaluate multiple pipeline configurations combining Named Entity Recognition (NER) models and large language models (LLMs). Performance was assessed by accuracy of maximum PI-RADS extraction and processing speed using self-hosted implementations of spaCy NER, Transformers NER, and generative LLMs Llama 3, Qwen3, and Gemma3. Results: Across the top 10 pipeline configurations, accuracy for maximum PI-RADS extraction ranged from 89.3% to 95.5%, with processing times spanning 150 milliseconds to 70 seconds per report. Generative LLM pipelines achieved the highest accuracy (up to 95.5%) but were substantially slower (2 to 70 seconds), whereas NER based pipelines demonstrated lower accuracy (88.5%) with faster performance (50 to 150 milliseconds). Discussion: Discriminative NER pipelines achieved high accuracy while offering advantages in speed and potential scalability. Accuracy gains from LLMs were accompanied by significantly higher computational cost, potentially limiting feasibility in high-volume clinical environments. Conclusion: Discriminative methods were more efficient than generative models in annotating PIRADS from mpMRI report text, providing insights into configurations for optimal clinical deployment when volume is a limiting factor. However, generative AI offered improved accuracy with less upfront development.

11
Reducing Under-Triage Risk in Large Language Model Based Clinical Triage Using UMLS-CUI Augmentation

Gokhale, R.; Kukreja, M.; Kumar, N.; Gourab, K.

2026-08-10 health informatics 10.64898/2026.08.07.26358932 medRxiv
Top 0.1%
6.9%
Show abstract

Background: Public facing large language models (LLMs) are increasingly used for health guidance, including triage recommendations. We evaluated whether augmenting LLM prompts with standardized clinical concepts from the Unified Medical Language System (UMLS) could improve the safety and robustness of clinical triage recommendations. Methods: We used a publicly available dataset comprising 60 clinician-authored clinical vignettes, each represented in 16 demographic and narrative variations, yielding 960 vignette-factor combinations. Clinical entities were extracted using a two-stage pipeline combining ClinicalBERT-based named entity recognition with rule-based identification of laboratory abnormalities. Extracted entities were mapped to UMLS Concept Unique Identifiers (CUIs). Negated concepts were excluded. A confidence-weighted CUI voting classifier was trained using empirical associations between CUIs and clinician-assigned triage categories. We compared five approaches: CUI-only classification, MedGemma 27B, MedGemma 27B augmented with CUIs, GPT-4o-mini, and GPT-4o-mini augmented with CUIs. Outcomes included overall accuracy, under-triage, over-triage, emergency-case accuracy, and sensitivity to anchoring statements. Results: CUI augmentation decreased under-triage but increased over-triage in both models tested (GPT-4o-mini and MedGemma 27B). It improved high-acuity recognition while reducing recognition of low-acuity cases. CUI augmentation had mixed effects on overall triage accuracy; accuracy increased for MedGemma 27B but decreased for GPT-4o-mini. Emergency-case accuracy improved from 73.0% to 80.7% for GPT-4o-mini and from 60.5% to 68.5% for MedGemma 27B. CUI augmentation also reduced susceptibility to anchoring statements. These findings suggest that the principal value of CUI augmentation may be shifting model behavior toward safety-oriented behavior rather than uniformly improving overall accuracy. Conclusion: Ontology-grounded prompt augmentation shifted LLM triage recommendations toward greater sensitivity to high-acuity presentations and reduced overall under-triage. These safety gains were accompanied by increased over-triage and mixed effects on overall accuracy. A hybrid architecture combining LLM-based language understanding with interpretable UMLS-derived clinical concepts may improve the safety and robustness of AI-assisted triage. Further evaluation using real-world patient communications and clinical outcomes is warranted.

12
Evaluating Eight Retrieval-Augmented Generation (RAG) Large Language Models' Responses to Clinical Questions: A Comparative Study

Krump, P. A.; Blasingame, M. N.; Koonce, T. Y.; Williams, A. M.; Su, J.; Giuse, N. B.

2026-08-12 health informatics 10.64898/2026.08.10.26360108 medRxiv
Top 0.1%
6.7%
Show abstract

Background: Large language models (LLMs) that use retrieval-augmented generation (RAG) are increasingly used to answer clinical questions, although the evaluation of these systems remains limited. Building on previous studies conducted by our team, this case report aimed to improve upon this knowledge gap by applying a reusable methodology to compare the performance of eight LLMs that utilize RAG techniques for evidence synthesis. Case Presentation: Eight commercially available RAG LLM tools (OpenEvidence, Undermind, Consensus, SciSpace, Elicit, MediSearch, EvidenceHunt, and Scite) were evaluated using twelve ChatGPT-generated clinical questions on the topics of treatment, etiology, and prognosis. To enable comparison, we prompted ChatGPT to identify all key unique medical concepts from the full set of LLM responses to each question. Concepts were categorized as critical ("must-have") or non-critical ("nice-to-have") for answering the clinical question. Experienced information scientists were consulted at each step for their expertise. Descriptive statistics and Kruskal-Wallis tests were used to compare performance across tools and question categories. No significant differences were found among the eight RAG LLMs in their coverage of "must-have" (p=0.95) or "nice-to-have" (p=0.16) key unique medical concepts, and no single tool consistently captured all identified concepts. Conclusions: These findings suggest that RAG LLMs may be supplementary tools for evidence retrieval and synthesis but cannot, at this time, fully replace comprehensive expert review of the medical literature. The evaluation framework presented here may be a useful model for future comparative assessments of rapidly evolving AI evidence synthesis tools.

13
Beyond Length of Stay: Patient and Carer Perspectives on Virtual Hospital Pathways Following Colorectal Surgery

Reza, L.; Arbai, Z.; Ward, H.; Payne, L.; Kinross, J.; Patel, V.

2026-08-27 surgery 10.64898/2026.08.24.26361282 medRxiv
Top 0.1%
6.6%
Show abstract

Background Virtual hospital (VH) pathways support early discharge through remote monitoring, but limited evidence has hindered implementation in colorectal surgery. This study aimed to define patient- and carer-relevant outcomes and experiences of VH following colorectal surgery. Methodology A patient and public involvement and engagement (PPIE) consultation was conducted with 8 participants (7 patients, 1 carer; 4 women, 4 men) who had experienced VH following bowel resection at a high-volume robotic unit. Purposive sampling ensured that 50% of participants had experienced readmission. The 90-minute session was delivered via Microsoft Teams. Data were analysed using reflexive thematic analysis. Results Seven themes were identified: readmission, remote monitoring, carer burden, recovery, equity, readiness for discharge, and information delivery. Patients supported early discharge when remote monitoring enabled timely detection of complications and readmission pathways were efficient. Readmission was not perceived as failure but as appropriate escalation. Dissatisfaction with readmission was related to delays in emergency care. Remote monitoring provided psychological safety, with patients feeling held at home. Carers assumed substantial, often unrecognised, quasi-clinical roles. Recovery was defined by return to function rather than length of stay. Equity concerns were evident, with VH favouring those with adequate support at home, digital literacy, and language proficiency. Discharge readiness was both clinical and psychological. Information delivery at discharge was often poorly retained and requires reinforcement preoperatively at every encounter with patients and carers. Conclusions VH pathways are acceptable and valued. Readmission is a marker of system responsiveness rather than failure of early discharge on VH. Psychological preparedness, carer support, and equitable access are critical to successful and scalable implementation of early discharge using a virtual hospital.

14
Pragmatic trial design of a digital supportive care platform for patients with brain tumours and their carers

Kalla, M.; Bray, S. C.; Schadewaldt, V.; Krishnasamy, M.; Whittle, J. R.; Chapman, W.; Huckvale, K.; Burns, K.; Capurro, D.; Layton, M. J.; Thomas, J.; Lourenco, R. D. A.; Andrew, D.; McAlpine, H.; Dhillon, R. S.; Cain, S.; Rosenthal, M.; Drummond, K. J.

2026-08-21 health informatics 10.64898/2026.08.18.26360754 medRxiv
Top 0.1%
6.6%
Show abstract

Patients with a brain tumour receive evidence-based clinical care in Australia but a focus on supportive care, including social connection, is often deficient. Digital health platforms hold promise to support these patients and their carers. Existing platforms often lack end-user co-design, evidence-based development and rigorous evaluation. Recognising this unmet need, we co-designed Brain Tumours Online, a digital supportive care platform to streamline access to educational resources, symptom management tools, and peer support for patients, carers, and healthcare professionals. In this article, we present our evaluation approach for Brain Tumours Online to advance methodological thinking in the evaluation of multi-faceted, co-designed digital health platforms. In contrast to standardised procedures in clinical trials, digital health interventions such as supportive care platforms are more complex due to their interactive nature, no prescriptive protocols for usage and the dynamic content of web-based information. Thus, traditional evaluation approaches often fall short in evaluating such multi-faceted digital health supportive care platforms. To address these challenges, we developed a bespoke, logic-modelling based evaluation approach to assess the usability, engagement, impact, and economic value of our platform. Our pragmatic but rigourous evaluation approach required the adaptation of existing evaluation frameworks, subject-matter, and lived experience expert knowledge. Our implementation science and co-design approach are shared in different papers. Our study outcomes will also be shared in a separate paper. In the current paper, we share our approach to the evaluation of Brain Tumours Online and provide insights that may be of value for other researchers interested in the nuances of trialing multi-faceted digital health supportive care platforms.

15
Design tensions in a two-sided marketplace for reusable digital therapeutics software components: a qualitative interview study

Kowatsch, T.; Melamed, S.; Nissen, M.; Merz, Y.

2026-07-20 health informatics 10.64898/2026.07.17.26358332 medRxiv
Top 0.1%
6.6%
Show abstract

Objectives To identify stakeholder-perceived design tensions in a two-sided marketplace for reusable digital therapeutics (DTx) software components and to use these tensions to propose alternative marketplace concepts. Methods We conducted 24 semi-structured interviews with digital health researchers and professionals. Data were analysed using hybrid deductive-inductive codebook thematic analysis. The Magic Triangle provided the initial deductive structure. One researcher coded all transcripts; a second independently applied the developing codebook to five transcripts to refine definitions and consistency. Seventeen parent themes were synthesized into 12 design tensions, which informed three author-generated marketplace concepts. Results Participants described trade-offs concerning target users and host, component scope and customization, quality labels, verification, geographic scope, pricing, interoperability, platform launch, risks and market niche. The resulting concepts emphasized a regional startup ecosystem, a research-oriented hybrid marketplace or a global marketplace with stricter entry requirements. Discussion The concepts combine the tensions in different ways and highlight competing priorities in governance, openness, assurance, scalability and early platform growth. Conclusion Stakeholders identified recurring design choices for a DTx software-component marketplace. The concepts provide hypotheses for prototyping and evaluation; the study did not test technical feasibility, market demand, regulatory acceptability or effects on development cost or time.

16
Uncertainty-aware extraction of clinical findings from Finnish EHRs using open large language models

Leinonen, J. V.; Knuutila, J.; Kurki, S.; Pamilo, S.; Koskinen, M.

2026-07-09 health informatics 10.64898/2026.07.07.26355248 medRxiv
Top 0.1%
6.6%
Show abstract

Objective. To evaluate whether open-weight large language models (LLMs) can accurately extract clinical findings from Finnish-language pediatric records, and whether prediction uncertainty can be used to triage cases for expert review to minimize manual work. Materials and Methods. Retrospective cohort of 97 pediatric ischaemic stroke patients (1 month - 17 years) from Helsinki University Hospital (2010 - 2023). Three open LLMs (gpt-oss-20b, DeepSeek-R1-Distill-Qwen-32B, and medgemma-27b-text-it) were prompted in English to detect four extraction targets (hemiplegia, headache, seizure, and stroke as a positive control) from each patient's full free-text record. Each combination received 15 calls (five temperatures x three repeats). Performance was benchmarked against a clinician reference (accuracy, recall, precision, F1). Shannon entropy across the 15 calls quantified within-model uncertainty; inter-model disagreement provided an ensemble signal. Patients were ranked by uncertainty for a simulated selective-review workflow. Findings were externally validated in an independent neonatal stroke cohort (n = 88). Results. Gpt-oss-20b achieved the best balance of recall (0.91 - 1.00) and precision (0.83 - 0.92), with F1 0.89 - 0.95 across non-control extraction targets. Entropy in misclassified cases was 2.4 - 3.4 times higher than in correctly classified cases. Entropy-based triage achieved complete error coverage by reviewing <10% of patients for hemiplegia (8.3%) and headache (8.2%), and 19.6% for seizure. Neonatal validation reached F1 0.95 for Apgar 1 min and binary seizure, and F1 0.87 for 4-class stroke-subtype classification. Discussion. Within-model entropy and inter-model disagreement provided complementary, calibrated signals of likely error in a non-English clinical setting. Conclusion. Open LLMs can extract clinical findings from Finnish pediatric records with accuracy comparable to published English benchmarks, and uncertainty-based triage substantially reduces required expert workload.

17
Evaluation of Large Language Models for Post-Cystectomy Sexual Health Counseling in Women: A Pilot Study

Shafau, F.; Dave, A. A.; Omole, I.; Guzman, T.; Rehman, N.; Enemchukwu, E.; Bresler, L.

2026-07-08 urology 10.64898/2026.06.25.26356154 medRxiv
Top 0.2%
5.8%
Show abstract

Abstract Objective To evaluate the adherence to guidelines and readability of large language model-generated sexual health information related to female sexual dysfunction following cystectomy, and to determine whether adherence differs across models and prompt formats. A secondary objective was to introduce an analytic strategy using principal component analysis to examine the dimensions of readability metrics. Methods Three large language models (LLMs), ChatGPT, Gemini, and Perplexity were prompted with six clinical questions related to sexual function after cystectomy. Questions were phrased in long-form and short-form language. Responses were independently graded by two reviewers, derived from guideline recommendations. Linear mixed-effects models predicted adherence as functions of LLM, prompt, and reviewer, with clinical questions as a random intercept. Readability was assessed using five metrics, and principal component analysis (PCA) was used to determine latent structure. Results ChatGPT demonstrated the highest (estimated marginal mean [emm] = 0.769), outperforming Gemini (0.499) and Perplexity (0.457). Shorter, less complex prompts elicited higher adherence than more complex, clinical prompts. All models produced content that exceeded recommended reading levels. PCA demonstrated that a single dominant component accounted for 76.7% of variance across readability indices, indicating a shared underlying construct. Conclusion ChatGPT produced the most guideline-concordant information overall. High linguistic complexity was seen across models, highlighting a barrier to patient comprehension. These findings characterize large language models as variable medical information systems whose outputs rely heavily on prompt structure and model type.

18
Accuracy and error patterns of ChatGPT-4o for real-time English-Nepali voice translation: A cross-sectional field evaluation in rural Nepal

Mandich, A.; Koirala, S.; Westen, S.; Adhikari, S.; Acharya, A.; Shrestha, A.

2026-08-28 health informatics 10.64898/2026.08.25.26361303 medRxiv
Top 0.2%
5.7%
Show abstract

Language discordance can impede community-based research and health communication where trained interpreters are limited. Although multimodal artificial intelligence systems can provide real-time spoken translation, performance with under-resourced languages during spontaneous field interactions remains poorly characterized. We evaluated ChatGPT-4o during bidirectional English-Nepali voice translation in a community setting near Dhulikhel Hospital, Nepal. In this cross-sectional field study, 30 primarily Nepali-speaking adults were recruited by convenience sampling. ChatGPT-4o mediated conversations using standardized English questions and spontaneous Nepali responses. A bilingual Nepali-English reviewer assessed 485 translated utterances using a 3-point accuracy scale and an inductively developed framework for translation and conversational deviations. Of 485 translations, 282 (58.1%) received the highest accuracy rating, 134 (27.6%) a moderate rating, and 69 (14.2%) the lowest. Mean accuracy was higher for English-to-Nepali than Nepali-to-English translation (2.63 {+/-} 0.53 vs 2.23 {+/-} 0.86); 63 of 69 low-accuracy translations (91.3%) occurred in the Nepali-to-English direction. Among 329 deviation tags, the most frequent were distortion of intended meaning (17.1%), overly formal or unnatural phrasing (14.7%), omission (14.2%), and addition of content (11.5%). Some fluent outputs substantially altered meaning or introduced information not expressed by the speaker. ChatGPT-4o demonstrated potential for real-time English-Nepali communication but also produced errors that could alter interpretation of participant responses. Accuracy was lower and more variable for Nepali-to-English translation; however, translation direction was confounded with input type because Nepali inputs were spontaneous and English inputs standardized, limiting conclusions about directional performance. These findings support cautious use for low-stakes conversational exchange and human verification when errors could affect research validity, clinical decisions, or participant understanding. As multimodal AI evolves, performance should be reevaluated across languages, real-world conditions, and model versions, with bilingual oversight and community partnership remaining central to responsible use.

19
Supporting people to access social security payments through the Special Rules for End of Life: a qualitative study of the perspectives of patients, carers and health care professionals

Davies, J. M.; Marshall, S.; Hussain, J.; Diggle, M.; French, M.; Stone, J.; Fimister, G.; Ogden, M.; Sleeman, K. E.; Bradshaw, A.; Harding, R. E.

2026-06-15 palliative medicine 10.64898/2026.06.12.26355509 medRxiv
Top 0.2%
5.7%
Show abstract

Background: People living with terminal illness face a double financial burden from additional costs and loss of earning for themselves and their carers. Social security benefits are intended to help alleviate some of this financial pressure, and in the UK and other countries people are eligible for fast-tracked access to financial support via the Special Rules for End of Life. One in 3 people who are eligible miss out on this support, yet there is limited evidence on the reasons for this take-up deficit. Objectives: The aim of this study is to understand the barriers and facilitators to claiming benefits for terminally ill people from the perspectives of patients, carers, and health care professionals. Methods: This is a qualitative study combining i) focus groups with healthcare professionals recruited via professional networks and social media, and ii) interviews with patients and carers recruited in hospital and hospice settings. We analysed the data using Practical Thematic Analysis Results: Fifty-five multidisciplinary healthcare professionals participated in 11 focus groups, and we interviewed 10 patients and carers. We constructed five descriptive themes to summarise the data: Navigating priorities and uncertainty; positive impacts alongside a sense of shame and stigma; talking about money, difficulties and dividends; everybodys, yet nobodys, responsibility; and sticking points in the system. Conclusion: The themes reveal several challenges that may contribute to people not taking up this financial support. However, discussions about access to benefits were also seen as a core part of holistic care, a positive way to offer support and a gateway to other discussions about end-of-life care preferences and decisions. Recommendations for policy and practice include evaluating the adoption of a diagnostic rather than a prognostic eligibility criteria, integrating discussions about benefits into existing processes such as advance care planning, and improving education and support for clinicians.

20
Comparing Human and Large Language Model Responses to Patients Online Questions: Towards Multi-dimensional Patient-centered Support

Hussein, M. A.; Doshi, R.; He, L.; Reynolds, T.

2026-07-17 health informatics 10.64898/2026.07.15.26355314 medRxiv
Top 0.2%
5.6%
Show abstract

Patients and caregivers seek informational and emotional support throughout medical care, especially when interpreting unfamiliar laboratory test results. Although resources such as patient portals and online health communities (OHCs) help address questions, gaps remain. The emergence of large language models (LLMs) offers the potential to be a complementary source of support to assist patients and caregivers in understanding and using their test results. The objective of our study is to empirically compare LLM responses to patients online questions containing their laboratory test results to responses written by peers in an OHC. We compared the 519 peer replies to 122 laboratory test-related posts from an OHC to 488 responses generated from four LLMs using mixed computational and qualitative methods. LLMs frequently provided clear explanations of medical terminology and structured interpretations of numeric results but were longer and less readable. Peers offered more personalized, context-specific emotional support. Overall, LLMs have the potential to complement peer responses in OHCs, but require greater emotional depth, reasoning transparency, and alignment with community norms.